Papers by Vivek Varadarajan Sembium

2 papers
Rationale-Guided Distillation for E-Commerce Relevance Classification: Bridging Large Language Models and Lightweight Cross-Encoders (2025.coling-industry)

Copied to clipboard

Challenge: Large-scale e-commerce search systems typically follow a multi-step process to retrieve relevant products for a given query.
Approach: They propose a distillation approach that uses "rationales" generated by Large Language Models to guide smaller cross-encoder models.
Outcome: The proposed model achieves ROC-AUC improvements of 1.4% on 9 multilingual e-commerce datasets, 2.4% on 3 ESCI datasets and 6% on GLUE datasets while being 50 times faster per sample.
Multilingual Continual Learning using Attention Distillation (2025.coling-industry)

Copied to clipboard

Challenge: Existing models for Query-product relevance classification are not accurate across multiple languages.
Approach: They propose a multilingual continual learning framework that adds adapters for each new language and incorporates a fusion layer above language-specific adapters.
Outcome: The proposed approach reduces trainable parameters by 80% while outperforming SOTA CL methods on proprietary and external datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations